Papers with content moderation strategies

4 papers
MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations (2025.emnlp-main)

Copied to clipboard

Challenge: generative language agents predict user behaviors such as liking, sharing, and flagging content.
Approach: They propose a framework where generative language agents predict user behaviors such as liking, sharing, and flagging content.
Outcome: The proposed framework analyzes content moderation strategies and user engagement dynamics at scale and demonstrates that agents’ articulated reasoning for their social interactions aligns with their collective engagement patterns.
Conspiracy Theories and Where to Find Them on TikTok (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on TikTok's potential to promote and amplify harmful content have not been conducted.
Approach: They analyze a longitudinal dataset of 1.5M videos shared in the U.S. over three years and evaluate the effects of TikTok’s Creativity Program for monetization.
Outcome: The proposed model achieves high precision in detecting harmful content, but its overall performance is comparable to fine-tuned traditional models such as RoBERTa.
Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated Tweets (2025.coling-main)

Copied to clipboard

Challenge: generative AI has been used to generate fluent and convincing text on social media platforms . a new study examines the generative capabilities of four popular large language models .
Approach: They propose a methodology to examine the generative capabilities of four prominent LLMs on Twitter using a dataset from Llama 3, Mistral, Qwen2 and GPT4o.
Outcome: The proposed method examines the generative capabilities of four prominent LLMs on Twitter.
FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes (2026.findings-acl)

Copied to clipboard

Challenge: Suicide memes are increasingly common on social media, yet remain poorly understood and potentially harmful.
Approach: They propose a dataset designed for fine-grained analysis of suicide memes and benchmark 16 models for figurative language, suicide severity, and content detection.
Outcome: The proposed model outperforms existing models on figurative language, suicide severity, and suicide-related content detection tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations